Skip to content

fix(opencode): allow governed free models for private repositories - #830

Open
seonghobae wants to merge 33 commits into
mainfrom
fix/opencode-private-free-opt-in-20260808
Open

fix(opencode): allow governed free models for private repositories#830
seonghobae wants to merge 33 commits into
mainfrom
fix/opencode-private-free-opt-in-20260808

Conversation

@seonghobae

@seonghobae seonghobae commented Aug 8, 2026

Copy link
Copy Markdown
Contributor

Problem and RCA

The central OpenCode review path historically used repository visibility as the proxy for anonymous/free-model eligibility. That is too coarse: a private repository can be intentionally public-equivalent, while absence of Actions secrets does not prove that tracked source, history, comments, fixtures, or generated review evidence are non-confidential.

A fresh exact-head review also exposed three distinct fail-closed defects in the first implementation: a private caller could bypass the immutable-base policy by pre-populating opencode-free/*; git ls-tree -z output was reconstructed rather than requiring the real terminating NUL; and the static free alias list had drifted from the current OpenCode Zen zero-cost catalog. These are source defects and are being repaired on this branch rather than classified as reviewer-capacity or governance blockers.

Solution

  • Add a fail-closed trusted-base policy at .github/opencode-private-free-models.json.

  • Require the exact canonical declaration:

    {
      "schema_version": 1,
      "allow_private_free_models": true,
      "repository_data_classification": "public_equivalent",
      "external_model_data_use_accepted": true
    }
  • Read the declaration from the exact immutable PR base commit and reject a head that adds, removes, renames, chmods, or modifies its own policy. The opt-in takes effect only for a later PR after normal protected-base integration.

  • Never treat preconfigured opencode-free/* text as authorization. Private or unverified callers have every anonymous candidate removed before policy evaluation; only the immutable base policy may re-enable them.

  • Preserve public behavior only from positive visibility evidence: an explicit trusted OPENCODE_REPOSITORY_IS_PRIVATE=false, or a credential-free successful Git read from a strictly validated public ContextualWisdomLab origin. Ambiguous, private, auth-required, malformed, or unavailable visibility evidence fails closed to the policy path.

  • Synchronize the governed anonymous pool to the currently documented OpenCode Zen zero-cost aliases: nemotron-3-ultra-free, deepseek-v4-flash-free, north-mini-code-free, laguna-s-2.1-free, ling-3.0-flash-free, big-pickle, and mimo-v2.5-free. Stale opencode-free/* aliases are filtered before provider execution.

  • Scope every OpenCode subprocess to its selected provider credential. Anonymous/free and export execution receives no GitHub token, Actions OIDC/runtime/cache/results credential, NVIDIA/OpenAI/OpenRouter/OpenCode key, or unrelated provider secret.

  • Recognize both long and short OpenCode model selectors (--model, --model=, -m, -m=), reject duplicate/missing selectors, and stop parsing at --.

  • Require the trusted policy tree lookup to contain exactly one real NUL-terminated git ls-tree -z record; truncated or extra records fail closed.

  • Validate integer model-pool runtime/retry/cycle controls before Bash arithmetic or timeout consumption, with reviewed safe defaults.

Current governed free catalog

The wrapper allowlist is intentionally narrower than the generated provider configuration. The current primary Zen documentation identifies these seven zero-cost aliases:

  1. opencode-free/nemotron-3-ultra-free
  2. opencode-free/deepseek-v4-flash-free
  3. opencode-free/north-mini-code-free
  4. opencode-free/laguna-s-2.1-free
  5. opencode-free/ling-3.0-flash-free
  6. opencode-free/big-pickle
  7. opencode-free/mimo-v2.5-free

Aliases previously labelled free for Hy3, MiniMax M3, GLM 5, Kimi K2.5, and Qwen3.6 Plus are not in the current documented zero-cost list and are no longer admitted by the wrapper. Catalog changes require a separately reviewable source change; the opencode-free/* prefix alone is never trusted as pricing evidence.

Security and governance properties

The canonical declaration means the repository owner accepts external free-model processing for tracked repository content classified as public_equivalent; it does not claim that secret scanning proves absence of confidential facts. Secret Protection, push protection, generic/custom patterns, and CODEOWNERS remain defense in depth.

The policy checker accepts only a regular non-executable 100644 blob at the fixed path, strict UTF-8, at most 4,096 bytes, exact field types/values, and JSON without duplicate keys. It accepts only full 40-character base/head SHAs, ignores user/system Git configuration, disables hooks/filesystem monitors, and fails closed on missing, invalid, changed, malformed-tree, or unreadable policy state.

Automated reviewer verdicts remain separate from the repository's qualifying counted independent human approval requirement. Reviewer rate limits or missing counted approval are governance/capacity evidence, not source defects and must not trigger speculative source patches.

Test-first repair after current-head review

The review-triggered repair was implemented test-first:

  • fail-closed private preconfigured-free and unknown-visibility contracts;
  • current official free-catalog filtering contract;
  • real ls-tree -z NUL termination and extra-record rejection;
  • -m/-m= aliases, -- terminator, duplicate/missing model selector, and executable-boundary tests;
  • integer runtime-control contracts for the delegated model pool.

Every new head invalidates predecessor-head checks/reviews. Current exact-head machine evidence is still regenerating after these repairs; queued/in-progress checks are not acceptance.

Operational acceptance

Code-level checks are necessary but not sufficient. Issue #833 remains the post-merge operational contract: separately merge the canonical policy into an authoritatively classified private low-risk canary, use a later PR to prove inherited-base activation, verify an actual opencode-free/* selection and zero credential exposure, run a private negative control without policy, preserve keyed fallback/exhaustion fail-closed behavior, and demonstrate or deterministically rehearse rollback. If no private repository can be authoritatively classified as public-equivalent, keep the feature inactive rather than inventing eligibility.

Migration note

This focused change supersedes the overlapping private/free-model routing slice in draft PR #760. Any future rebase or decomposition of #760 must preserve this trusted-base opt-in, catalog filtering, visibility fail-closed path, and provider-scoped credential boundary rather than restoring blanket private exclusion or trusting candidate text.

Sources

Summary by CodeRabbit

  • 새로운 기능

    • 비공개 저장소에서 익명 무료 모델을 사용할 수 있는 엄격한 정책 검증을 추가했습니다.
    • 공개 동등 데이터 선언과 외부 모델 데이터 사용 동의가 확인된 경우에만 무료 모델을 활성화합니다.
    • 무료 모델 후보를 검증된 목록으로 제한하고, 오래되거나 알 수 없는 모델은 제외합니다.
    • 모델 제공자별 자격 증명을 격리하고, 익명·알 수 없는 모델에는 민감한 자격 증명을 전달하지 않습니다.
  • 문서

    • 비공개 무료 모델 사용 조건과 운영 절차를 문서화했습니다.
  • 테스트

    • 정책 검증, 모델 선택, 자격 증명 보호 및 실패 시 안전한 차단 동작을 검증하는 테스트를 추가했습니다.

Add an immutable trusted-base opt-in for private repositories classified as public-equivalent, preserve the existing model-pool implementation byte-for-byte behind a policy wrapper, and isolate every OpenCode subprocess to its selected provider credential.
@coderabbitai

coderabbitai Bot commented Aug 8, 2026

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Private 저장소의 익명 OpenCode 무료 모델 사용을 base 커밋 정책으로 제한했습니다. 모델 풀을 별도 구현으로 위임하고, 공급자별 자격 증명 격리, 실행 제어, fail-closed 테스트를 추가했습니다.

Changes

OpenCode 거버넌스 및 실행 제어

Layer / File(s) Summary
Private 무료 모델 정책 검증
scripts/ci/opencode_private_free_model_policy.py, docs/doctoring/..., docs/examples/..., CHANGELOG.md, tests/test_opencode_private_free_model_policy_*.py
Base 커밋의 고정 정책 blob만 평가합니다. 정책 변경, 잘못된 SHA와 Git tree, 손상된 JSON, 중복 키, 잘못된 UTF-8 및 canonical 값은 거부합니다.
모델 풀 위임 및 무료 후보 제어
scripts/ci/run_opencode_review_model_pool.sh, tests/test_opencode_private_free_model_runner_contract.py, tests/test_opencode_delegated_runner_contract.py
Wrapper가 구현 계약과 정책 검사기를 확인한 뒤 sibling 구현을 실행합니다. 검증된 base 정책이 있을 때만 익명 무료 후보를 중복 없이 추가합니다.
공급자 자격 증명 격리
scripts/ci/opencode_provider_guard.sh, tests/test_opencode_provider_guard.py
선택된 공급자의 자격 증명만 OpenCode에 전달합니다. GitHub, Actions OIDC 및 미선택 공급자 키를 제거합니다. 익명 모델, export, 알 수 없는 공급자는 자격 증명 없이 실행합니다.
위임된 모델 풀 실행
scripts/ci/run_opencode_review_model_pool_impl.sh
모델 출력 검증, 승인 게이트, 공급자 오류 분류, 재시도, 백오프, 후보 격리, 실행 시간·시도·예산 제한 및 exhausted 상태 기록을 구현합니다.

Estimated code review effort: 4 (Complex) | ~75 minutes

Possibly related issues

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
  participant Wrapper as run_opencode_review_model_pool.sh
  participant Policy as opencode_private_free_model_policy.py
  participant Guard as opencode_provider_guard.sh
  participant Pool as run_opencode_review_model_pool_impl.sh
  participant OpenCode as OpenCode
  Wrapper->>Policy: base/head 커밋으로 정책 평가
  Policy-->>Wrapper: 무료 모델 사용 허용 또는 거부
  Wrapper->>Guard: OpenCode 실행 wrapper 설치
  Wrapper->>Pool: 후보 목록과 실행 환경 전달
  Pool->>Guard: 선택된 모델 실행 요청
  Guard->>OpenCode: 정리된 자격 증명 환경으로 실행
  OpenCode-->>Pool: 모델 출력과 세션 결과 반환
  Pool-->>Wrapper: 성공, 재시도 또는 exhausted 상태 기록
Loading
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 69.53% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed 제목은 private repository에서 governed free model을 허용하는 PR의 주요 변경 사항을 정확하고 간결하게 설명합니다.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/opencode-private-free-opt-in-20260808

Comment @coderabbitai help to get the list of available commands.

Expand the governed private pool to every anonymous candidate already configured by the central workflow and retain the established fail-closed source contract while delegating runtime behavior to the unchanged implementation.
Keep the established central source-level contract visible at the stable entrypoint and fail closed on a truncated delegated implementation in full workflow materializations.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🧹 Nitpick comments (6)
tests/test_opencode_private_free_model_runner_contract.py (1)

206-221: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

이 테스트는 두 개의 독립된 차단 이유를 동시에 만족합니다.

base_has_policy=False이므로 정책 평가가 이미 거부됩니다. 동시에 후보 목록에 opencode-free/glm-5-free가 있어 candidate_list_contains_anonymous_free_model이 조기 반환합니다. 따라서 "기존 free 풀은 재정렬하지 않는다"는 계약이 단독으로 검증되지 않습니다. base_has_policy=True로 바꾸면 조기 반환 경로만 검증합니다.

💚 테스트 강화 제안
     source, base_sha, head_sha = create_source_repository(
         tmp_path,
-        base_has_policy=False,
+        base_has_policy=True,
         head_changes_policy=False,
     )
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/test_opencode_private_free_model_runner_contract.py` around lines 206 -
221, Update test_existing_public_free_pool_is_not_reordered_or_duplicated to set
base_has_policy=True while keeping head_changes_policy=False, so the policy gate
passes and the test isolates the existing public free-pool ordering behavior
without triggering the anonymous free-model early return.
scripts/ci/run_opencode_review_model_pool.sh (2)

169-170: 🩺 Stability & Availability | 🔵 Trivial | ⚡ Quick win

trapinstall_provider_guard 앞에 등록하십시오.

현재 trap cleanup_provider_guard EXIT INT TERMinstall_provider_guard 다음 줄에 있습니다. mktemp -d 성공 후 cp 또는 chmod가 실패하면 set -e가 스크립트를 종료합니다. 그 시점에는 trap이 아직 없으므로 임시 디렉터리가 남습니다. trap을 먼저 등록하면 모든 실패 경로에서 정리가 실행됩니다.

♻️ 순서 변경 제안
-install_provider_guard
 trap cleanup_provider_guard EXIT INT TERM
+install_provider_guard
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/ci/run_opencode_review_model_pool.sh` around lines 169 - 170,
Register the cleanup trap before calling install_provider_guard so
cleanup_provider_guard handles failures during temporary-directory setup,
including cp or chmod errors under set -e.

77-103: 🩺 Stability & Availability | 🔵 Trivial | 💤 Low value

후보 목록 확장 시 glob 확장을 차단하십시오.

for candidate in ${OPENCODE_MODEL_CANDIDATES:-}는 인용을 생략하여 단어 분리를 의도합니다. 그러나 파일명 확장도 함께 활성화됩니다. 워크플로가 *, ?, [를 포함한 후보 문자열을 전달하면 후보 이름이 현재 디렉터리 파일명으로 치환될 수 있습니다. 두 함수를 set -f/set +f로 감싸거나, read -r -a로 배열을 만들면 확장이 차단됩니다.

🛡️ 제안
 candidate_list_contains_anonymous_free_model() {
-  local candidate
-  for candidate in ${OPENCODE_MODEL_CANDIDATES:-}; do
+  local candidate
+  local -a candidates
+  read -r -a candidates <<<"${OPENCODE_MODEL_CANDIDATES:-}"
+  for candidate in "${candidates[@]}"; do
     case "$candidate" in
       opencode-free/*)
         return 0
         ;;
     esac
   done
   return 1
 }
 
 prepend_unique_anonymous_free_candidates() {
   local combined=""
   local candidate
-  for candidate in $anonymous_free_candidates ${OPENCODE_MODEL_CANDIDATES:-}; do
+  local -a candidates
+  read -r -a candidates <<<"$anonymous_free_candidates ${OPENCODE_MODEL_CANDIDATES:-}"
+  for candidate in "${candidates[@]}"; do
     case " $combined " in
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/ci/run_opencode_review_model_pool.sh` around lines 77 - 103, Disable
pathname expansion while iterating over OPENCODE_MODEL_CANDIDATES in
candidate_list_contains_anonymous_free_model and
prepend_unique_anonymous_free_candidates, preserving intentional
whitespace-based word splitting. Restore the caller’s globbing state after each
function completes, including early returns, or use a read-based array approach
that prevents glob expansion without changing candidate parsing.
scripts/ci/opencode_private_free_model_policy.py (1)

176-177: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

blob SHA 검증에 커밋 SHA 패턴을 재사용합니다.

COMMIT_SHA_PATTERN은 40자 16진수만 허용합니다. SHA-256 오브젝트 포맷 저장소에서 git ls-tree는 64자 SHA를 반환합니다. 그 경우 정책 평가는 상태 2로 실패합니다. 현재 GitHub 호스팅 저장소는 SHA-1이므로 즉시 영향은 없습니다. 별도의 오브젝트 ID 패턴(40 또는 64자)을 사용하면 향후 마이그레이션에서 안전합니다.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/ci/opencode_private_free_model_policy.py` around lines 176 - 177,
Update the validation around entry.object_sha in the policy evaluation flow to
use a dedicated object ID pattern that accepts valid 40- or 64-character
hexadecimal SHAs, rather than COMMIT_SHA_PATTERN. Keep the existing
PolicyEvaluationError and invalid-SHA handling unchanged.
scripts/ci/run_opencode_review_model_pool_impl.sh (1)

42-51: 🩺 Stability & Availability | 🔵 Trivial | 💤 Low value

normalize_opencode_output은 호출 컨텍스트의 errexit 비활성화에 의존합니다.

set -euo pipefail이 활성 상태입니다. Line 44의 opencode_review_approve_gate.sh가 0이 아닌 상태로 끝나면, 조건 컨텍스트 밖에서는 errexit이 발동하여 Line 46의 rc=$?와 Line 50의 rm -f "$probe"가 실행되지 않습니다. 현재 유일한 호출 지점인 Line 536은 if ! 조건이므로 동작합니다. 향후 다른 위치에서 호출하면 임시 파일이 남고 폴백이 중단됩니다. 명시적으로 상태를 잡으면 호출 위치와 무관하게 안전합니다.

♻️ 제안
 	if python3 "$GITHUB_WORKSPACE/scripts/ci/opencode_review_normalize_output.py" \
 		"$HEAD_SHA" "$RUN_ID" "$RUN_ATTEMPT" "$probe"; then
-		bash "$GITHUB_WORKSPACE/scripts/ci/opencode_review_approve_gate.sh" \
-			"$HEAD_SHA" "$RUN_ID" "$RUN_ATTEMPT" "$probe" >/dev/null
-		rc=$?
+		rc=0
+		bash "$GITHUB_WORKSPACE/scripts/ci/opencode_review_approve_gate.sh" \
+			"$HEAD_SHA" "$RUN_ID" "$RUN_ATTEMPT" "$probe" >/dev/null || rc=$?
 	else
 		rc=1
 	fi
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/ci/run_opencode_review_model_pool_impl.sh` around lines 42 - 51,
Update the normalize_opencode_output flow around opencode_review_approve_gate.sh
so its nonzero status is captured explicitly without relying on an outer if or !
condition to suppress errexit. Ensure rc is assigned before cleanup, rm -f
"$probe" always runs, and the function returns the captured status for callers
regardless of invocation context.
tests/test_opencode_private_free_model_policy_1.py (1)

17-113: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

세 테스트 파일이 동일한 97줄 헤더를 복제합니다. 공유 헬퍼 모듈이 없어 모듈 로더, run, git, commit_all, write_policy, repository fixture, evaluate가 세 번 정의되었습니다. 정책 검사기 인터페이스가 바뀌면 세 곳을 모두 수정해야 합니다. tests/conftest.py 또는 전용 헬퍼 모듈로 추출하십시오.

  • tests/test_opencode_private_free_model_policy_1.py#L17-L113: 헬퍼와 fixture를 공유 모듈로 옮기고 import로 대체하십시오.
  • tests/test_opencode_private_free_model_policy_2.py#L17-L113: 동일한 공유 모듈을 import하도록 바꾸십시오.
  • tests/test_opencode_private_free_model_policy_3.py#L17-L113: 동일한 공유 모듈을 import하도록 바꾸십시오.

참고: 세 파일 모두 sys.modules["opencode_private_free_model_policy"]에 서로 다른 모듈 객체를 등록합니다. 공유 모듈로 통합하면 이 중복 등록도 사라집니다.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/test_opencode_private_free_model_policy_1.py` around lines 17 - 113,
Extract the duplicated module loader, run, git, commit_all, write_policy,
repository fixture, and evaluate helpers into one shared test helper module.
Update tests/test_opencode_private_free_model_policy_1.py#L17-L113,
tests/test_opencode_private_free_model_policy_2.py#L17-L113, and
tests/test_opencode_private_free_model_policy_3.py#L17-L113 to import the shared
helpers and remove their local definitions, including separate sys.modules
registrations for opencode_private_free_model_policy.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/doctoring/opencode-private-free-model-policy.md`:
- Around line 63-74: 문서의 모델 목록에서 1번, 3번, 12번 항목의 잘린 `-fre` 접미사를 `-free`로 수정해
`anonymous_free_candidates` 및 `EXPECTED_FREE_CANDIDATES`와 이름을 일치시키세요.

In `@scripts/ci/opencode_private_free_model_policy.py`:
- Around line 218-219: Update the validation around EXPECTED_POLICY to compare
JSON values with strict type sensitivity, so boolean true is not accepted as
numeric 1 and vice versa. Preserve the exact canonical declaration requirement
for every field, including schema_version and allow_private_free_models, while
retaining the existing PolicyDenied behavior for mismatches.

In `@scripts/ci/opencode_provider_guard.sh`:
- Around line 16-24: Update the argument scan around previous_argument and
model_candidate to recognize both “--model candidate” and “--model=candidate”
forms. Track occurrences explicitly and reject duplicate --model values before
provider credential removal, while preserving the existing candidate validation
and single-model behavior.

In `@scripts/ci/run_opencode_review_model_pool_impl.sh`:
- Line 463: Validate OPENCODE_FATAL_ERROR_POLL_SECONDS through the existing
env_integer_or_default helper when assigning fatal_poll_seconds, preserving the
default of 5 for unset or non-integer values. Ensure the validated value is used
by the sleep call in the kill -0 polling loop.

---

Nitpick comments:
In `@scripts/ci/opencode_private_free_model_policy.py`:
- Around line 176-177: Update the validation around entry.object_sha in the
policy evaluation flow to use a dedicated object ID pattern that accepts valid
40- or 64-character hexadecimal SHAs, rather than COMMIT_SHA_PATTERN. Keep the
existing PolicyEvaluationError and invalid-SHA handling unchanged.

In `@scripts/ci/run_opencode_review_model_pool_impl.sh`:
- Around line 42-51: Update the normalize_opencode_output flow around
opencode_review_approve_gate.sh so its nonzero status is captured explicitly
without relying on an outer if or ! condition to suppress errexit. Ensure rc is
assigned before cleanup, rm -f "$probe" always runs, and the function returns
the captured status for callers regardless of invocation context.

In `@scripts/ci/run_opencode_review_model_pool.sh`:
- Around line 169-170: Register the cleanup trap before calling
install_provider_guard so cleanup_provider_guard handles failures during
temporary-directory setup, including cp or chmod errors under set -e.
- Around line 77-103: Disable pathname expansion while iterating over
OPENCODE_MODEL_CANDIDATES in candidate_list_contains_anonymous_free_model and
prepend_unique_anonymous_free_candidates, preserving intentional
whitespace-based word splitting. Restore the caller’s globbing state after each
function completes, including early returns, or use a read-based array approach
that prevents glob expansion without changing candidate parsing.

In `@tests/test_opencode_private_free_model_policy_1.py`:
- Around line 17-113: Extract the duplicated module loader, run, git,
commit_all, write_policy, repository fixture, and evaluate helpers into one
shared test helper module. Update
tests/test_opencode_private_free_model_policy_1.py#L17-L113,
tests/test_opencode_private_free_model_policy_2.py#L17-L113, and
tests/test_opencode_private_free_model_policy_3.py#L17-L113 to import the shared
helpers and remove their local definitions, including separate sys.modules
registrations for opencode_private_free_model_policy.

In `@tests/test_opencode_private_free_model_runner_contract.py`:
- Around line 206-221: Update
test_existing_public_free_pool_is_not_reordered_or_duplicated to set
base_has_policy=True while keeping head_changes_policy=False, so the policy gate
passes and the test isolates the existing public free-pool ordering behavior
without triggering the anonymous free-model early return.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4ef09089-44f5-44d9-997c-fbf050bce76d

📥 Commits

Reviewing files that changed from the base of the PR and between 6eb06cd and 3362860.

📒 Files selected for processing (13)
  • CHANGELOG.md
  • docs/doctoring/opencode-private-free-model-policy.md
  • docs/examples/opencode-private-free-models.json
  • scripts/ci/opencode_private_free_model_policy.py
  • scripts/ci/opencode_provider_guard.sh
  • scripts/ci/run_opencode_review_model_pool.sh
  • scripts/ci/run_opencode_review_model_pool_impl.sh
  • tests/test_opencode_delegated_runner_contract.py
  • tests/test_opencode_private_free_model_policy_1.py
  • tests/test_opencode_private_free_model_policy_2.py
  • tests/test_opencode_private_free_model_policy_3.py
  • tests/test_opencode_private_free_model_runner_contract.py
  • tests/test_opencode_provider_guard.py

Comment thread docs/doctoring/opencode-private-free-model-policy.md Outdated
Comment thread scripts/ci/opencode_private_free_model_policy.py Outdated
Comment thread scripts/ci/opencode_provider_guard.sh
Comment thread scripts/ci/run_opencode_review_model_pool_impl.sh Outdated
Expose every established central shell-gate marker at the stable model-pool entrypoint and verify the delegated implementation before any model process starts.
@seonghobae
seonghobae enabled auto-merge (squash) August 8, 2026 09:59
@opencode-agent
opencode-agent Bot disabled auto-merge August 8, 2026 10:00

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/test_opencode_private_free_model_runner_contract.py`:
- Around line 275-280: In test_wrapper_preserves_every_quick_gate_runner_marker,
correct the undefined wrapper_tex reference in the marker assertion loop to use
the existing wrapper_text variable read from WRAPPER.
- Line 116: Update the test fixture’s candidate-selection logic to use
OPENCODE_MODEL_CANDIDATES exclusively, remove the unused
OPENCODE_MODD_CANDIDATES fallback, and capture/assert that the first candidate
is selected because the fake opencode does not validate --model. Also correct
the undefined wrapper_tex reference to wrapper_text.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 42094baf-07a5-4706-8dcb-044a8071e605

📥 Commits

Reviewing files that changed from the base of the PR and between 3362860 and 23eb9f7.

📒 Files selected for processing (2)
  • scripts/ci/run_opencode_review_model_pool.sh
  • tests/test_opencode_private_free_model_runner_contract.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • scripts/ci/run_opencode_review_model_pool.sh

Comment thread tests/test_opencode_private_free_model_runner_contract.py Outdated
Comment thread tests/test_opencode_private_free_model_runner_contract.py Outdated
Correct the exact-head regression test so the stable wrapper marker contract is evaluated without truncating the final identifier.
Restore the previously verified exact runner-contract test blob while retaining the live quick-gate assertions in the central shell gate itself.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/test_opencode_private_free_model_runner_contract.py`:
- Line 237: 손상된 테스트 코드를 복구해 `for script in (...)` 구문이 `WRAPPER`와
`PROVIDER_GUARD`를 순회하도록 수정하고, 각 스크립트에 대해 Bash `-n` 구문 검증을 수행하게 하십시오. 변경 후 전체 테스트
스위트를 실행해 검증하십시오.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 0581ad70-cdc7-4c1b-90cf-b3c384e52202

📥 Commits

Reviewing files that changed from the base of the PR and between 23eb9f7 and 58de30c.

📒 Files selected for processing (1)
  • tests/test_opencode_private_free_model_runner_contract.py

Comment thread tests/test_opencode_private_free_model_runner_contract.py Outdated

Copy link
Copy Markdown
Contributor Author

@opencode-agent address

Exact-head bounded GREEN repair for e427cfb14dfc8ff0e678b25ea44662b3be3f74b3; target implementation blob scripts/ci/run_opencode_review_model_pool_impl.sh = 986982e9af3e65cf468f993d8e858d9e1edfc5c1. Abort without writing if either identity moved.

The current branch already contains RED contracts in tests/test_opencode_private_free_model_runner_contract.py for two still-valid CodeRabbit findings. Make only the minimal implementation changes needed to satisfy them:

  1. Assign fatal_poll_seconds through the existing env_integer_or_default OPENCODE_FATAL_ERROR_POLL_SECONDS 5 helper so malformed environment input cannot make sleep fail inside the set +e polling loop and busy-spin the runner.
  2. In normalize_opencode_output, capture a nonzero opencode_review_approve_gate.sh status explicitly (for example rc=0; ... || rc=$?) so cleanup of the temporary probe always executes regardless of invocation context; return the captured status unchanged.

Do not alter model selection, credentials, review semantics, timeout policy beyond this validation, or any unrelated file. Run the focused runner contract and the repository-authoritative exact-head suite before committing. Do not merge or synthesize approval.

Copy link
Copy Markdown
Contributor Author

@opencode-agent address

Exact-head bounded GREEN repair for e427cfb14dfc8ff0e678b25ea44662b3be3f74b3; scripts/ci/run_opencode_review_model_pool_impl.sh blob is 986982e9af3e65cf468f993d8e858d9e1edfc5c1. Do not write if either identity has moved.

Current-head Strix Changed Path Quality CI run 31253529097, job 93093433574, has one deterministic failure after the rest of the suite: 1 failed, 1026 passed, 16 subtests passed. The fail-first contract is tests/test_opencode_private_free_model_runner_contract.py::test_delegated_runner_validates_poll_interval_and_cleans_normalization_probe.

Two production defects are directly evidenced in the delegated stable implementation:

  1. run_one_model_attempt() currently uses fatal_poll_seconds="${OPENCODE_FATAL_ERROR_POLL_SECONDS:-5}". Invalid/empty/non-numeric user configuration can reach sleep or create a zero-second busy loop. Use the existing trusted helper exactly as the RED contract requires:
    fatal_poll_seconds="$(env_integer_or_default OPENCODE_FATAL_ERROR_POLL_SECONDS 5)".

  2. normalize_opencode_output() calls the approval gate under set -e and assigns rc=$? on the following line. A nonzero approval-gate result can therefore terminate before rm -f "$probe", leaking the temporary normalization probe. Preserve cleanup by initializing/capturing status with an explicit non-errexit conditional, matching the existing RED contract: the gate invocation must end with >/dev/null || rc=$?, then remove the probe and return the captured status. Do not turn a gate failure into success.

Make only these minimal production repairs plus any strictly necessary test/doc wording alignment. Do not weaken the fail-first assertions, do not change provider eligibility, credentials, wrapper policy, model order, secrets, workflows, or branch protection, and do not create temporary/self-modifying workflows.

Run the focused runner contract first, then the complete repository tests, bash scripts/ci/test_strix_quick_gate.sh, Bash syntax, Python compile/docstring/coverage gates, and exact-head security gates. Commit normally to this existing branch; do not merge or synthesize approval.

Copy link
Copy Markdown
Contributor Author

@opencode-agent address

Same unchanged exact head e427cfb14dfc8ff0e678b25ea44662b3be3f74b3: include the remaining unresolved current-head CodeRabbit documentation finding in the same bounded repair. In docs/doctoring/opencode-private-free-model-policy.md, correct only the three truncated candidate IDs so the operator runbook matches the executable/test contracts exactly:

  • opencode-free/nemotron-3-ultra-freopencode-free/nemotron-3-ultra-free
  • opencode-free/north-mini-code-freopencode-free/north-mini-code-free
  • opencode-free/qwen3.6-plus-freopencode-free/qwen3.6-plus-free

Do not alter candidate ordering or any other model identifier. Validate the documentation against anonymous_free_candidates and EXPECTED_FREE_CANDIDATES as part of the same exact-head run. Resolve the review thread only after the committed text is verified.

Copy link
Copy Markdown
Contributor Author

@opencode-agent address

Repair only the exact current head e427cfb14dfc8ff0e678b25ea44662b3be3f74b3 on protected base 6eb06cdd08c79a06f7b390069d4ffa49e2eb7dba. Immediately before any write, refetch the PR head/base and scripts/ci/run_opencode_review_model_pool_impl.sh; the exact current implementation blob is 986982e9af3e65cf468f993d8e858d9e1edfc5c1. If any identity moved, do not write and re-plan from the new state.

Exact-head Strix run 31253529097, job 93093433574, passed 1,026 tests and failed only test_delegated_runner_validates_poll_interval_and_cleans_normalization_probe. The permanent test already exposes two real fail-closed defects in the delegated runner; keep the test unchanged and make only these minimal production fixes:

  1. Replace the unvalidated fatal_poll_seconds="${OPENCODE_FATAL_ERROR_POLL_SECONDS:-5}" with fatal_poll_seconds="$(env_integer_or_default OPENCODE_FATAL_ERROR_POLL_SECONDS 5)", preventing malformed/empty hostile configuration from reaching sleep or creating a busy/failing poll loop.
  2. In normalize_opencode_output, preserve cleanup when opencode_review_approve_gate.sh returns nonzero under set -e by using the existing test-required form "$HEAD_SHA" "$RUN_ID" "$RUN_ATTEMPT" "$probe" >/dev/null || rc=$? (initialize/retain rc so rm -f "$probe" always executes before returning the gate status).

Do not change the trusted-base private/free-model policy, provider credential isolation, model list/order, reviewer identities or credentials, NVIDIA NIM behavior, branch protection, tests, or unrelated model-pool semantics. Do not add a temporary/write-capable repair workflow. Run the focused private-free-model runner contract first, then the complete central suite and Strix exact-head gate. Keep the branch unmerged; any new head requires fresh exact-head security/review evidence.

seonghobae commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

@opencode-agent address

Repair only the exact current-head deterministic/review blockers on PR #830. Live head is e427cfb14dfc8ff0e678b25ea44662b3be3f74b3, protected base is 6eb06cdd08c79a06f7b390069d4ffa49e2eb7dba, scripts/ci/run_opencode_review_model_pool_impl.sh blob is 986982e9af3e65cf468f993d8e858d9e1edfc5c1, and docs/doctoring/opencode-private-free-model-policy.md blob is 59693a35485f70f414961db2922d9a499331d2ae. Abort without writing if any head/base/target-blob identity moves.

Exact-head Strix Changed Path Quality CI run 31253529097, job 93093433574, completed 1,026 tests plus 16 subtests with exactly one reported failure at tests/test_opencode_private_free_model_runner_contract.py::test_delegated_runner_validates_poll_interval_and_cleans_normalization_probe. Two current CodeRabbit threads on this same head are also valid: the fatal poll interval is unvalidated, and doctoring truncates three governed model IDs.

Make the smallest repair, limited to the delegated implementation, its focused regression only if needed to cover zero, and the doctoring list:

  1. In run_one_model_attempt, replace raw ${OPENCODE_FATAL_ERROR_POLL_SECONDS:-5} consumption with fatal_poll_seconds="$(env_integer_or_default OPENCODE_FATAL_ERROR_POLL_SECONDS 5)", then fail safe from zero as well ([ "$fatal_poll_seconds" -gt 0 ] || fatal_poll_seconds=5) so malformed, negative, or zero configuration cannot create a no-delay watcher loop. Preserve the existing sleep call and default 5 seconds. If the current regression does not cover zero, add one bounded focused assertion/test before the production change; do not weaken the existing RED assertion.
  2. In normalize_opencode_output, capture approval-gate rejection without relying on caller if ! context: set rc=0 before invoking opencode_review_approve_gate.sh, run the gate as ... "$probe" >/dev/null || rc=$?, then keep unconditional rm -f "$probe" and return "$rc". Do not convert normalization failure or gate rejection into success.
  3. In docs/doctoring/opencode-private-free-model-policy.md, correct only the three truncated governed IDs: nemotron-3-ultra-frenemotron-3-ultra-free, north-mini-code-frenorth-mini-code-free, and qwen3.6-plus-freqwen3.6-plus-free.

Do not modify wrapper governance, provider credential isolation, candidate ordering, workflows, permissions, CHANGELOG, unrelated tests, or any other file. Run the focused runner contract, relevant delegated-runner tests, Bash syntax, and doc/model-name consistency before committing, then let normal exact-head CI/security/Strix and review gates rerun. Do not mark Ready/merge/release, resolve unrelated threads, or synthesize approval.

Copy link
Copy Markdown
Contributor Author

@opencode-agent

Perform a read-only formal review of exact current head 9a9b3e061599c28905aed793fc803123a2615205 against protected base 6eb06cdd08c79a06f7b390069d4ffa49e2eb7dba for PR #830. Re-fetch both identities before review and do not write, address, merge, or synthesize approval if either moved.

RCA focus: the former blanket private-repository exclusion conflated repository visibility with data classification. Verify that the replacement is operationally realistic and fail-closed: anonymous opencode-free/* candidates activate only from an unchanged trusted-base canonical public_equivalent declaration; a PR cannot self-enable; anonymous/free and export subprocesses receive no provider secrets, GitHub token, Actions OIDC/runtime/cache/results credentials, or unrelated provider keys; private repositories without the trusted-base opt-in remain on the private-safe keyed path.

All exact-head GitHub Actions checks currently report success and all inline review threads are resolved/outdated. Review the current source and bounded evidence independently. Submit a formal APPROVE only if affirmative source-backed evidence supports the contract; otherwise submit source-anchored REQUEST_CHANGES. Do not treat secret-scan success alone as proof that repository content is public-equivalent.

@opencode-agent
opencode-agent Bot disabled auto-merge August 8, 2026 18:07

Copy link
Copy Markdown
Contributor Author

@opencode-agent address

Repair only PR #830 if it is still exact head 6cf29abc43e6a1891ea05de0a4283d8d5bf83098 on base 6eb06cdd08c79a06f7b390069d4ffa49e2eb7dba. Immediately before any write, refetch the PR head/base and these target blobs; abort without writing if any identity moved:

  • tests/test_opencode_model_pool_runner.py blob 08d17f000859dbc36d77784bc04f79c790c817d9
  • tests/test_opencode_agent_contract.py blob daeaa37a25d5c76db946321820c028444194b85d

RCA from exact-head Strix run 31271895841, job 93139195742: the complete suite has exactly four failures (4 failed, 1036 passed, 16 subtests passed). These are compatibility-test failures caused by the newly fail-closed private/free wrapper architecture, not evidence that production privacy filtering should be weakened.

  1. Three failures in tests/test_opencode_model_pool_runner.py (test_free_provider_runtime_cap_preserves_queue_budget, test_nvidia_nim_combined_budget_preserves_fallback_attempt, test_free_provider_gets_one_bounded_schema_repair_attempt) use run_failed_model(), which creates a non-git source directory and supplies opencode-free/* candidates without any trusted visibility signal. The current wrapper correctly classifies unknown visibility as private/unverified and removes those candidates before immutable-base policy approval. These tests are intended to exercise generic free-provider runtime behavior, not the private-policy gate. Smallest realistic fix: make that test helper explicitly declare its fixture public with OPENCODE_REPOSITORY_IS_PRIVATE=false, preserving all production fail-closed behavior and leaving the dedicated private/unknown-visibility tests authoritative.

  2. test_workflow_provisions_sandbox_tool_and_reviewer_agent reads only scripts/ci/run_opencode_review_model_pool.sh but still asserts historical implementation markers such as OPENCODE_TOTAL_RETRY_BUDGET_SECONDS:-1500. The production architecture now intentionally delegates the stable pool implementation to run_opencode_review_model_pool_impl.sh and verifies its markers before execution. Smallest realistic fix: make the test inspect the composed wrapper + delegated implementation contract (for example concatenate both file texts) rather than adding inert compatibility literals/comments to production code.

Reject these alternatives: do not weaken unknown/private filtering, do not re-enable anonymous candidates without trusted public/base-policy evidence, do not add fake production marker comments solely to satisfy the test, do not change model catalog/order, credentials, policy checker, provider guard, workflows, branch protection, or reviewer authority.

Run the four failing tests first, then the complete repository suite and bash scripts/ci/test_strix_quick_gate.sh, Python compilation, Bash syntax, git diff --check, and all exact-head gates. Commit only the minimal test compatibility changes. Do not approve, merge, rebase, force-push, create temporary/write-capable workflows, or synthesize review evidence.

@opencode-agent opencode-agent Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

OpenCode cannot approve yet because required coverage evidence did not pass.

Review outcome

1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence

  • Problem: The required coverage-evidence job result was failure, so OpenCode cannot establish approval sufficiency for this head.

  • Root cause: Automated approval is only valid when the same-head coverage-evidence job proves supported repository test suites passed and configured docstring gates passed or were advisory, or reports not applicable because no supported source files or package manifests exist. Missing, failed, skipped, unavailable, or unsupported-tooling test evidence is a blocker.

  • Fix: Install or configure the repository test/docstring evidence tooling when source files or package manifests exist, rerun the current-head coverage-evidence job, and approve only after it reports success with required evidence or explicit no-source not-applicable evidence.

  • Regression test: Keep the approval branch checking needs.coverage-evidence.result == success before posting APPROVE, and publish REQUEST_CHANGES when coverage-evidence blocker states such as cancelled, skipped, failed, unsupported-tooling, or below-100 evidence are present.

  • Result: REQUEST_CHANGES

  • Reason: coverage-evidence result was failure, so required test/docstring evidence was not proven for current head 6cf29abc43e6a1891ea05de0a4283d8d5bf83098.

  • Head SHA: 6cf29abc43e6a1891ea05de0a4283d8d5bf83098

  • Workflow run: 31273012550

  • Workflow attempt: 1

Coverage evidence

Coverage Decision

  • Result: FAIL
  • Test evidence: not proven passing
  • Docstring evidence: not proven passing when configured
  • Failure count: 1

Changed-File Evidence Map

flowchart LR
  PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
  Evidence --> S1["Changed file: CHANGELOG.md"]
  S1 --> I1["repository behavior"]
  I1 --> R1["Review risk: Changed file: CHANGELOG.md"]
  R1 --> V1["required checks"]
  Evidence --> S2["Docs (2 files)"]
  S2 --> I2["operator or user guidance"]
  I2 --> R2["Review risk: Docs (2 files)"]
  R2 --> V2["docs review"]
  Evidence --> S3["CI script (4 files)"]
  S3 --> I3["review and security gate shell path"]
  I3 --> R3["Review risk: CI script (4 files)"]
  R3 --> V3["bash -n plus Strix self-test"]
  Evidence --> S4["Test (6 files)"]
  S4 --> I4["regression suite"]
  I4 --> R4["Review risk: Test (6 files)"]
  R4 --> V4["targeted test run"]
Loading

@opencode-agent

opencode-agent Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor

OpenCode Review Overview

  • Head SHA: a18cd1a7a61dd5abaad8cf9deebcf3245158ac69
  • Workflow run: 31284309359
  • Workflow attempt: 1
  • Gate result: REQUEST_CHANGES (approval step)

Pull request overview

OpenCode cannot approve yet because required coverage evidence did not pass.

Review outcome

1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence

  • Problem: The required coverage-evidence job result was failure, so OpenCode cannot establish approval sufficiency for this head.

  • Root cause: Automated approval is only valid when the same-head coverage-evidence job proves supported repository test suites passed and configured docstring gates passed or were advisory, or reports not applicable because no supported source files or package manifests exist. Missing, failed, skipped, unavailable, or unsupported-tooling test evidence is a blocker.

  • Fix: Install or configure the repository test/docstring evidence tooling when source files or package manifests exist, rerun the current-head coverage-evidence job, and approve only after it reports success with required evidence or explicit no-source not-applicable evidence.

  • Regression test: Keep the approval branch checking needs.coverage-evidence.result == success before posting APPROVE, and publish REQUEST_CHANGES when coverage-evidence blocker states such as cancelled, skipped, failed, unsupported-tooling, or below-100 evidence are present.

  • Result: REQUEST_CHANGES

  • Reason: coverage-evidence result was failure, so required test/docstring evidence was not proven for current head a18cd1a7a61dd5abaad8cf9deebcf3245158ac69.

  • Head SHA: a18cd1a7a61dd5abaad8cf9deebcf3245158ac69

  • Workflow run: 31284309359

  • Workflow attempt: 1

Coverage evidence

Coverage Decision

  • Result: FAIL
  • Test evidence: not proven passing
  • Docstring evidence: not proven passing when configured
  • Failure count: 1

Changed-File Evidence Map

flowchart LR
  PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
  Evidence --> S1["Changed file: CHANGELOG.md"]
  S1 --> I1["repository behavior"]
  I1 --> R1["Review risk: Changed file: CHANGELOG.md"]
  R1 --> V1["required checks"]
  Evidence --> S2["Docs (2 files)"]
  S2 --> I2["operator or user guidance"]
  I2 --> R2["Review risk: Docs (2 files)"]
  R2 --> V2["docs review"]
  Evidence --> S3["CI script (4 files)"]
  S3 --> I3["review and security gate shell path"]
  I3 --> R3["Review risk: CI script (4 files)"]
  R3 --> V3["bash -n plus Strix self-test"]
  Evidence --> S4["Test (8 files)"]
  S4 --> I4["regression suite"]
  I4 --> R4["Review risk: Test (8 files)"]
  R4 --> V4["targeted test run"]
Loading

Copy link
Copy Markdown
Contributor Author

@opencode-agent address

Exact-current-head RCA repair only. Re-fetch PR #830 immediately before writing and abort if any identity below moved:

  • head 6cf29abc43e6a1891ea05de0a4283d8d5bf83098
  • base 6eb06cdd08c79a06f7b390069d4ffa49e2eb7dba
  • tests/test_opencode_agent_contract.py blob daeaa37a25d5c76db946321820c028444194b85d
  • tests/test_opencode_model_pool_runner.py blob 08d17f000859dbc36d77784bc04f79c790c817d9
  • governance wrapper scripts/ci/run_opencode_review_model_pool.sh blob aa0df2db540ecc71cdbd56bd6ad0a735b5454a4a
  • focused governance contract tests/test_opencode_private_free_model_runner_contract.py blob 148399666e4624c0d8ba3a5c51678feb42f78bc7

RCA from exact-head Strix run 31271895841, job 93139195742: the complete suite reached 4 failed, 1036 passed, 16 subtests passed. The failures are stale test ownership after the model-pool split, not a production private/free-policy failure:

  1. test_workflow_provisions_sandbox_tool_and_reviewer_agent still binds model_pool_runner to the governance wrapper and then asserts low-level implementation markers such as OPENCODE_TOTAL_RETRY_BUDGET_SECONDS:-1500. Those markers are now owned by run_opencode_review_model_pool_impl.sh; the wrapper correctly verifies/delegates to that sibling and owns visibility/policy/catalog/credential-guard behavior.
  2. Three behavioral tests in test_opencode_model_pool_runner.py exercise free-provider timeout/schema-repair implementation behavior by invoking the governed wrapper with free-only candidates but no public-visibility or immutable-base policy authorization. The wrapper correctly filters those candidates, so the intended low-level implementation path is never reached.

Feasibility screening:

  • Do NOT weaken or bypass the wrapper's private/unverified free-candidate filtering; that would reintroduce the security defect fixed by this PR.
  • Do NOT duplicate implementation literals/functions back into the wrapper merely to satisfy stale assertions.
  • Do NOT fabricate a trusted public/private policy fixture in generic implementation tests unless the test is explicitly intended to exercise wrapper integration; governance already has focused coverage in test_opencode_private_free_model_runner_contract.py.
  • The smallest realistic remediation is test-only ownership alignment: low-level model-pool implementation assertions/behavioral tests should target scripts/ci/run_opencode_review_model_pool_impl.sh, while wrapper policy/visibility/catalog/provider-guard behavior remains exclusively covered by the focused governance contract.

Make only the minimal test-only changes necessary:

  • in tests/test_opencode_agent_contract.py, read the delegated implementation for model_pool_runner implementation-marker assertions; keep workflow assertions that the workflow invokes run_opencode_review_model_pool.sh, and preserve/add a narrow assertion that the wrapper delegates to/verifies run_opencode_review_model_pool_impl.sh if needed for clarity;
  • in tests/test_opencode_model_pool_runner.py, run implementation-behavior fixtures against run_opencode_review_model_pool_impl.sh rather than the governance wrapper;
  • do not modify production scripts, model catalog, credentials, permissions, branch protection, policy semantics, or unrelated tests.

Acceptance: first rerun the four exact failing tests, then the focused private-free governance contract, provider-guard tests, complete repository pytest/coverage suite, Bash syntax, and Strix quick gate. Commit only an ordinary descendant commit on the existing branch if all intended focused tests are GREEN. Do not mark Ready/merge/approve/release; every new head must regenerate exact-head CI/security/review evidence.

Copy link
Copy Markdown
Contributor Author

@opencode-agent address

Exact-current-head bounded test-compatibility repair only. Re-fetch PR #830 immediately before writing and abort without writing if any identity moved:

  • head 1ff4f65561149ea5a2553b8b9bc8b5026ec5f31b
  • base 6eb06cdd08c79a06f7b390069d4ffa49e2eb7dba
  • target tests/test_opencode_agent_contract.py blob daeaa37a25d5c76db946321820c028444194b85d

RCA from predecessor exact-head Strix run 31271895841 was 4 failed, 1036 passed, 16 subtests passed. One direct test-only repair has already landed on the current branch: run_failed_model() now explicitly marks its generic free-provider fixture public, preserving the production fail-closed private/unknown policy. The only remaining known compatibility failure is test_workflow_provisions_sandbox_tool_and_reviewer_agent: it reads only the wrapper text but asserts historical implementation literal OPENCODE_TOTAL_RETRY_BUDGET_SECONDS:-1500. The current architecture deliberately delegates the implementation, and run_opencode_review_model_pool_impl.sh now validates this control as env_integer_or_default OPENCODE_TOTAL_RETRY_BUDGET_SECONDS 1500.

Feasibility decision: do not add an inert old marker to production and do not weaken visibility/policy behavior. Make only the smallest test-side correction needed to assert the current delegated implementation contract. Prefer changing that stale assertion to the current env_integer_or_default OPENCODE_TOTAL_RETRY_BUDGET_SECONDS 1500 contract and, if necessary, make the test read the delegated implementation as well as the wrapper. Do not modify production scripts, model catalog/order, credentials, policy checker, provider guard, workflows, branch protection, or reviewer authority.

Run the known failing contract first, then the complete repository suite, Strix quick gate, Python compilation, Bash syntax, and git diff --check. Commit only the minimal test compatibility change as an ordinary descendant. Do not approve, merge, rebase, force-push, create a temporary/write-capable workflow, or reuse predecessor-head evidence.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/doctoring/opencode-private-free-model-policy.md`:
- Around line 227-228: 참고 문헌의 조회 날짜를 현재 실제 날짜인 August 8, 2026으로 수정하세요. OpenCode
Zen 참고 문헌의 URL과 나머지 인용 형식은 그대로 유지하세요.

In `@scripts/ci/opencode_provider_guard.sh`:
- Around line 21-24: The model-value handling in
scripts/ci/opencode_provider_guard.sh must reject missing, selector-like, and
terminator arguments: when expect_model_value is set, treat -- as a missing
value and treat --model, -m, --model=*, and -m=* as duplicate-selector errors
instead of consuming them as the model. Extend
tests/test_opencode_provider_guard.py with these three input forms, verifying
exit code 64 and that no child process runs.

In `@scripts/ci/run_opencode_review_model_pool.sh`:
- Around line 111-121: Reject zero for OPENCODE_POOL_CYCLE_SLEEP_SECONDS by
changing its minimum validation value from 0 to 1 in
scripts/ci/run_opencode_review_model_pool.sh lines 111-121. Also update the
direct execution path in scripts/ci/run_opencode_review_model_pool_impl.sh lines
796-808 so cycle_sleep values less than or equal to 0 are restored to 60 before
the model-pool cycle sleep is used.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: a06ada63-f56a-461e-bd82-7bccde1c5a0d

📥 Commits

Reviewing files that changed from the base of the PR and between 9a9b3e0 and 1ff4f65.

📒 Files selected for processing (9)
  • docs/doctoring/opencode-private-free-model-policy.md
  • scripts/ci/opencode_private_free_model_policy.py
  • scripts/ci/opencode_provider_guard.sh
  • scripts/ci/run_opencode_review_model_pool.sh
  • scripts/ci/run_opencode_review_model_pool_impl.sh
  • tests/test_opencode_model_pool_runner.py
  • tests/test_opencode_private_free_model_policy_1.py
  • tests/test_opencode_private_free_model_runner_contract.py
  • tests/test_opencode_provider_guard.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • scripts/ci/opencode_private_free_model_policy.py

Comment thread docs/doctoring/opencode-private-free-model-policy.md
Comment thread scripts/ci/opencode_provider_guard.sh
Comment thread scripts/ci/run_opencode_review_model_pool.sh

Copy link
Copy Markdown
Contributor Author

@opencode-agent address

Supersede the earlier exact-head test-only handoff in comment 5227905520 with this complete current-head bounded repair. Work only if PR #830 is still exact head 1ff4f65561149ea5a2553b8b9bc8b5026ec5f31b on base 6eb06cdd08c79a06f7b390069d4ffa49e2eb7dba. Immediately before any write, refetch the PR head/base and every target blob below; abort without writing if any identity moved:

  • tests/test_opencode_agent_contract.py = daeaa37a25d5c76db946321820c028444194b85d
  • scripts/ci/opencode_provider_guard.sh = b841f096cf028303a2749f7be641fcce5ebd3f47
  • tests/test_opencode_provider_guard.py = 94f91dff392fc97e2385cfbe79e68c08d062bb67
  • scripts/ci/run_opencode_review_model_pool.sh = aa0df2db540ecc71cdbd56bd6ad0a735b5454a4a
  • scripts/ci/run_opencode_review_model_pool_impl.sh = 05d8f80db25e03124b43a71cf1bceb09f68dd551
  • tests/test_opencode_model_pool_runner.py = 4a504ffd1ec96051efede2d87acad49b63ca1bc2

RCA from exact-head Strix run 31276125807, job 93149813421: the complete central suite reached 1039 passed, 16 subtests passed with exactly one deterministic failure in tests/test_opencode_agent_contract.py::test_workflow_provisions_sandbox_tool_and_reviewer_agent. That test reads only run_opencode_review_model_pool.sh, but the architecture now deliberately delegates execution behavior to run_opencode_review_model_pool_impl.sh; it still asserts implementation-only markers such as OPENCODE_TOTAL_RETRY_BUDGET_SECONDS:-1500 against the governance wrapper. This is stale test ownership, not a production model-pool regression.

Fresh CodeRabbit review on this same head also exposed two valid fail-closed source defects:

  1. opencode_provider_guard.sh: while waiting for the value after --model/-m, a second selector (--model, -m, --model=*, -m=*) or the -- terminator can be consumed as the model value rather than rejected. That can misclassify provider credential scope. Add RED cases first for --model -m=openai/gpt-5.4, --model --, and one repeated long-selector form. Require exit 64, empty stdout, and no child execution; then minimally reject selector-like/terminator values while preserving valid --model, --model=, -m, -m= and post--- parsing semantics.
  2. model-pool cycle cadence: OPENCODE_POOL_CYCLE_SLEEP_SECONDS=0 is currently accepted by the wrapper and direct implementation. With unbounded cycles/budget controls this permits immediate retry cycling; even under the default provider-attempt ceiling it can create an avoidable request burst. Add a RED contract first, then make the wrapper require minimum 1 second and make the delegated implementation independently restore cycle_sleep<=0 to the reviewed 60-second default before any cycle sleep. Preserve the existing deadline clamp: when a positive deadline leaves no sleep budget, terminate rather than forcing a 60-second sleep past the deadline.

Feasibility decision:

  • Fixing test ownership is the smallest realistic remedy for the Strix failure. Do NOT copy historical implementation literals back into the wrapper and do NOT collapse the reviewed wrapper/implementation boundary merely to satisfy a stale assertion.
  • The selector and zero-cycle findings are current exact-head source defects with deterministic local acceptance tests and require no new credential, permission, ruleset change, or provider assumption; fix them test-first.
  • The CodeRabbit retrieval-date comment is NOT part of this repair. The repository automation operates in Asia/Seoul, where the current local date is August 9, 2026; Retrieved August 9, 2026 is therefore not a future date. That thread has been resolved without changing the citation.

Implementation constraints:

  • In test_workflow_provisions_sandbox_tool_and_reviewer_agent, read both the governance wrapper and run_opencode_review_model_pool_impl.sh; keep wrapper-policy assertions on the wrapper, and bind implementation/runtime markers to the delegated implementation. Update the retry-budget assertion to the current validated implementation contract (env_integer_or_default OPENCODE_TOTAL_RETRY_BUDGET_SECONDS 1500) rather than resurrecting the obsolete raw parameter-expansion literal.
  • Do not change private/public eligibility, trusted-base policy semantics, the seven governed free aliases, provider credential sets, reviewer identities, permissions, model ordering, branch protection, or unrelated workflow behavior.
  • Do not create a temporary/self-modifying/write-capable repair workflow; use ordinary descendant commits only.

Acceptance order: run the newly added RED cases first; then the exact failing agent-contract test, complete provider-guard tests, model-pool runner tests, private-free governance contract, Bash syntax for wrapper/implementation/guard, Python compilation, complete repository pytest/coverage/docstring gates, and bash scripts/ci/test_strix_quick_gate.sh. Commit only after the intended focused RED cases are GREEN. Leave the PR unmerged and do not synthesize approval; every new head must regenerate exact-head CI/security/review evidence.

@opencode-agent opencode-agent Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

OpenCode cannot approve yet because required coverage evidence did not pass.

Review outcome

1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence

  • Problem: The required coverage-evidence job result was failure, so OpenCode cannot establish approval sufficiency for this head.

  • Root cause: Automated approval is only valid when the same-head coverage-evidence job proves supported repository test suites passed and configured docstring gates passed or were advisory, or reports not applicable because no supported source files or package manifests exist. Missing, failed, skipped, unavailable, or unsupported-tooling test evidence is a blocker.

  • Fix: Install or configure the repository test/docstring evidence tooling when source files or package manifests exist, rerun the current-head coverage-evidence job, and approve only after it reports success with required evidence or explicit no-source not-applicable evidence.

  • Regression test: Keep the approval branch checking needs.coverage-evidence.result == success before posting APPROVE, and publish REQUEST_CHANGES when coverage-evidence blocker states such as cancelled, skipped, failed, unsupported-tooling, or below-100 evidence are present.

  • Result: REQUEST_CHANGES

  • Reason: coverage-evidence result was failure, so required test/docstring evidence was not proven for current head a18cd1a7a61dd5abaad8cf9deebcf3245158ac69.

  • Head SHA: a18cd1a7a61dd5abaad8cf9deebcf3245158ac69

  • Workflow run: 31284309359

  • Workflow attempt: 1

Coverage evidence

Coverage Decision

  • Result: FAIL
  • Test evidence: not proven passing
  • Docstring evidence: not proven passing when configured
  • Failure count: 1

Changed-File Evidence Map

flowchart LR
  PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
  Evidence --> S1["Changed file: CHANGELOG.md"]
  S1 --> I1["repository behavior"]
  I1 --> R1["Review risk: Changed file: CHANGELOG.md"]
  R1 --> V1["required checks"]
  Evidence --> S2["Docs (2 files)"]
  S2 --> I2["operator or user guidance"]
  I2 --> R2["Review risk: Docs (2 files)"]
  R2 --> V2["docs review"]
  Evidence --> S3["CI script (4 files)"]
  S3 --> I3["review and security gate shell path"]
  I3 --> R3["Review risk: CI script (4 files)"]
  R3 --> V3["bash -n plus Strix self-test"]
  Evidence --> S4["Test (8 files)"]
  S4 --> I4["regression suite"]
  I4 --> R4["Review risk: Test (8 files)"]
  R4 --> V4["targeted test run"]
Loading

Copy link
Copy Markdown
Contributor Author

@opencode-agent address

Repair only PR #830 at exact current head a18cd1a7a61dd5abaad8cf9deebcf3245158ac69 on protected base 6eb06cdd08c79a06f7b390069d4ffa49e2eb7dba. Immediately before any write, refetch the PR head/base plus tests/test_opencode_agent_contract.py blob daeaa37a25d5c76db946321820c028444194b85d and scripts/ci/run_opencode_review_model_pool_impl.sh blob 5900ac1cfc923ac4d3ca1e90368c17c6648ea2e4. Abort without writing if any identity moved.

RCA: exact-head Strix run 31283698013, job 93168950013, completed 1044 passed, 16 subtests passed with exactly one failure: tests/test_opencode_agent_contract.py::test_workflow_provisions_sandbox_tool_and_reviewer_agent. The failing assertion still requires implementation-owned historical literal OPENCODE_TOTAL_RETRY_BUDGET_SECONDS:-1500 inside the governance wrapper run_opencode_review_model_pool.sh. Runtime ownership was delegated to run_opencode_review_model_pool_impl.sh, whose exact current implementation correctly uses budget_seconds="$(env_integer_or_default OPENCODE_TOTAL_RETRY_BUDGET_SECONDS 1500)". This is stale test ownership, not a production model-pool regression.

Feasibility: do not resurrect the obsolete literal in production, do not change wrapper governance, provider/catalog/credential logic, #624 retirement routing, workflows, branch protection, or any production file. The smallest verifiable remediation is test-only: keep the wrapper assertions that genuinely describe wrapper/governance behavior, read scripts/ci/run_opencode_review_model_pool_impl.sh in the same test (or otherwise bind this exact implementation-owned contract), and replace only the obsolete wrapper literal assertion with a positive assertion for the current delegated integer-normalization contract env_integer_or_default OPENCODE_TOTAL_RETRY_BUDGET_SECONDS 1500 in the implementation.

Run the focused failing test first, then the complete repository suite and Strix quick/quality contract appropriate to this branch. Commit an ordinary descendant commit only if the fail-first assertion becomes green without weakening any other test. Do not approve, merge, mark Draft, rebase, force-push, or synthesize review evidence.

Copy link
Copy Markdown
Contributor Author

@jules

Please make exactly one test-only ownership repair on PR #830, and nothing else. Before writing, re-fetch and require all of these identities to remain unchanged: PR head a18cd1a7a61dd5abaad8cf9deebcf3245158ac69, protected base 6eb06cdd08c79a06f7b390069d4ffa49e2eb7dba, and tests/test_opencode_agent_contract.py blob daeaa37a25d5c76db946321820c028444194b85d. Abort without writing if any moved.

RCA: Strix run 31283698013 has one deterministic failure at test_workflow_provisions_sandbox_tool_and_reviewer_agent: the test still asserts implementation-owned historical literal OPENCODE_TOTAL_RETRY_BUDGET_SECONDS:-1500 in governance wrapper scripts/ci/run_opencode_review_model_pool.sh, but runtime ownership was deliberately delegated to scripts/ci/run_opencode_review_model_pool_impl.sh, which normalizes the value through env_integer_or_default OPENCODE_TOTAL_RETRY_BUDGET_SECONDS 1500. Production behavior is correct; this is stale test ownership.

Smallest accepted change: edit only tests/test_opencode_agent_contract.py. In that existing test, read scripts/ci/run_opencode_review_model_pool_impl.sh separately and replace the obsolete positive wrapper-literal assertion with an assertion that the delegated implementation contains the normalized env_integer_or_default OPENCODE_TOTAL_RETRY_BUDGET_SECONDS 1500 contract. Also assert the obsolete raw OPENCODE_TOTAL_RETRY_BUDGET_SECONDS:-1500 literal is absent from the governance wrapper if that fits the existing test structure. Do not modify production scripts, workflows, model/provider lists, credentials, permissions, docs, or any other test. Do not skip/xfail or weaken the test.

Run the focused failing test first, then the complete repository suite and Strix/quality contracts available to you. Commit as an ordinary descendant only; no amend/rebase/force-push. Do not mark Ready, approve, merge, or alter branch protection. If any guard moved or a different writer becomes active, stop and report the live mismatch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant